01 The Big Picture — Five Axes to Watch
Model news is noisy. Axis-level news is durable. Almost every credible announcement of the recent period lands on one of five axes — and each axis maps to a hardware or software constraint this series has already derived.
02 The Frontier Scoreboard
Each trend × its mathematical essence × where the series derives it.
| Trend | One-line math essence | Series doc |
|---|---|---|
| Test-time compute (o1/R1 style) | Quality ∝ thinking tokens; tokens billed as output (3–5× input price) | Doc 18 · Test-Time Compute |
| MoE everywhere | 671B total / 37B active ≈ 5.5% of params per token | Doc 21 · SOTA LLMs & MoE |
| KV-cache innovation wave | KV bytes = 2 · layers · kv-heads · head-dim · seq · bytes; shrink any factor | Doc 08 · GPU Memory · Doc 22 · KV Cache Types |
| Alternative architectures | SSM state O(1) vs attention O(n²); 1.58-bit ternary weights | Doc 02 · Transformers · Doc 25 · New Architectures |
| Decision models (JEV) | Probabilities read from hidden states; state reuse Q·S → S tokens | Item 01 below |
| Agent standards & harnessing | Tool schemas are cached prefix tokens; harness = context manager | Doc 12 · Agents · 09 · Context Solutions · Doc 27 · Standardization |
| Speculative & prefill scheduling | E[accepted draft tokens] × cheap verify vs expensive decode | Doc 14 · Decoding · 10 · Inference Stack |
| On-device & small models | 4-bit weights: bytes/param ÷ 4; NPU SRAM beats HBM latency | Doc 10 · Inference Stack · Doc 20 · On-Device AI |
| Long-context battles | Attention O(n²) compute, O(n) KV memory — and fidelity decays mid-context | Doc 06 · Inference · 11 · RAG |
| Post-training everywhere | RLHF/DPO commoditized; RLVR rewards = checkable answers | Doc 17 · Post-Training |
03 One Anim — The 2026 Frontier Map
Step through the axes in the order they arrived. Each stop is one bet about where the marginal dollar of compute buys the most capability.
04 The Deep-Dive List
One mini-card per trend: what it is, why engineers care, the math where it's natural, and where the series derives it properly.
01Decision Models — JEV (TypeSafe)
What it is (as of this writing): JEV is reportedly a decision model — a model that outputs typed decisions with calibrated probabilities directly, e.g. Yes: 80% / No: 20%, instead of generating text and hoping the prose parses. The only detailed public source is a Medium piece by Bijit Ghosh, "Inside JEV: Architecture of a Decision Model" — the architecture is not officially disclosed, so treat everything below as a report about an emerging class, not a verified spec.
The reported mechanism: the architecture was reconstructed black-box from ~10k API calls. A causal transformer encodes the shared context once and reuses its KV/state across Q question branches — for context length S, the per-question state work reportedly drops from ~Q·S to S tokens. Each branch ("route this ticket" / "estimate urgency" / "flag for review") attends to the shared context plus only its own options; attention masks isolate sibling questions so they can't leak into each other. Probabilities are read directly from hidden states — no autoregressive decode loop at all. Training reportedly targets calibration (RLCD — Reinforcement Learning for Calibrated Decisions, plausibly via a log-loss or Brier-style objective), and the reported MMLU calibration error is ≈ 0.031, concentrated in high-confidence predictions.
Why engineers care: this is expected-cost routing as a first-class primitive. An agent harness can branch on a calibrated probability instead of regex-parsing prose — and the routing inequality turns a model output directly into an escalation budget. It slots into the agent-loop economics of doc 12 and the context-reuse strategy of doc 09: one encoded context, many decision branches, is exactly the cache-friendly shape doc 07 preaches.
02Test-Time Compute & Reasoning Models
What: o1/R1-style models emit thousands of hidden "thinking" tokens before the visible answer. Why it matters: thinking tokens are billed as output — the 3–5× price class from doc 07 — so accuracy is now a per-token purchase. RLVR (reinforcement learning with verifiable rewards) trains the thinking; budget forcing caps it. Math hook: E[answer quality] rises roughly with thinking-token budget, with diminishing returns — your job is finding the knee of that curve per task class. Deep dive: doc 18.
03MoE Everywhere
What: DeepSeek-V3 (publicly reported 671B total / 37B active parameters), Llama-4, Qwen3-MoE, GLM-4.5 — sparsity is now the default frontier recipe. Why it matters: capability scales with total parameters while per-token cost tracks active ones (≈5.5% here) — but only if routing keeps experts balanced, hence aux-loss-free balancing schemes, and only if memory fits, hence MLA. Math hook: cost per token ∝ active params, memory ∝ total params. Deep dive: dense-vs-sparse grounding in doc 02; the MoE doc is doc 21 · SOTA LLMs & MoE.
04The KV-Cache Innovation Wave
What: GQA became standard, MLA compresses K/V into low-rank latents, paged KV eliminates fragmentation, FP8/INT4 KV halves-or-quarters the bytes. Why it matters: the KV cache is the memory wall of doc 08 — KV bytes = 2 · layers · kv-heads · head-dim · seq · bytes-per-elem — and every term of that product is now an engineering target. Math hook: shrink any factor, shrink batchable concurrency proportionally. Deep dive: doc 08; the dedicated KV doc is doc 22 · KV Cache Types.
05Alternative Architectures
What: a wave of non-vanilla-attention designs, as of this writing mostly at the hybrid or niche stage:
| Architecture | Key math | What it trades |
|---|---|---|
| Mamba/SSM hybrids (Jamba, Zamba) | Recurrent state, O(n) in sequence length | Gives up exact recall at distance; gains linear cost |
| RWKV | Attention folded into RNN-style state updates | Constant-memory inference for weaker parallel scoring |
| Gated DeltaNet | Delta-rule state updates with gating | Compression loss vs full attention fidelity |
| BitNet b1.58 | Ternary weights {−1, 0, 1}: log₂3 ≈ 1.58 bits | Extreme efficiency for lower per-weight precision |
| Byte-latent transformers | Patches bytes adaptively, no tokenizer | Patching complexity for tokenizer robustness |
| Diffusion LMs (LLaDA, Mercury) | Parallel denoising, not left-to-right decode | Speed for changed sampling/edition semantics |
| JEPA world models | Predict in representation space, not token space | Abstraction for direct pixel/token prediction |
| Titan / test-time memory | Learned memory updated at inference time | Extra state machinery for long-horizon recall |
Why it matters: each row re-attacks the O(n²) attention or 16-bit weight assumptions derived in doc 02. None has displaced the transformer yet; all are worth tracking. Deep dive: doc 25 · New Architectures Landscape.
06Agent Standards & Harnessing
What: MCP (Model Context Protocol) as a tool-interoperability standard, plus harness patterns — orchestrator-workers, skills, memory banks. Why it matters: this is the "software" frontier of the series: the static + dynamic shift, where tool schemas become cached prefix tokens (doc 07) and the harness becomes a context manager deciding what earns a slot in the window. Math hook: every standardized tool schema is amortized cache input, not fresh input. Deep dives: doc 12, doc 09; the standards doc is doc 27 · Standardization.
07Speculative & Prefill Scheduling
What: draft-verify decoding (EAGLE, Medusa) uses a cheap drafter whose guesses a strong model verifies in parallel; chunked prefill interleaves prompt reading with decoding; disaggregated serving (Splitwise, DistServe, Mooncake) splits prefill and decode onto different machines. Why it matters: this is the serving-layer SLO handshake — trading verification compute for bandwidth-bound decode speed. Math hook: speedup ≈ E[tokens accepted per draft] × drafter-cheapness − verify overhead; chunked prefill trades TTFT against per-token latency (ITL). Deep dives: doc 14, doc 10; the serving doc is doc 24 · Prefill & Decoding Strategies.
08On-Device & Small Models
What: NPUs in consumer silicon, 4-bit quantization as a shipping default, WebGPU bringing inference to the browser. Why it matters: edge became a first-class substrate — private, offline, zero marginal API cost. Math hook: 4-bit weights cut the bytes-per-param term of the memory equation by 4×; NPU SRAM wins on latency where HBM bandwidth wins on throughput (the CPU-vs-GPU asymmetry of Perception · CPU vs GPU). Deep dive: doc 10 covers quantization; the edge doc is doc 20 · Small Models & On-Device AI.
09Long-Context Battles
What: a public arms race toward 1M–2M-token windows — versus the quieter math of attention fidelity. Why it matters: attention compute is O(n²) and KV memory is O(n) per token forever, and empirically retrieval fidelity sags mid-context ("lost in the middle"). Ring attention shards the n² across devices; RAG keeps n small and retrieves. Math hook: a 2M-token window doesn't repeal doc 06's per-decode-step KV read — long context permanently taxes output speed. Deep dives: doc 06, doc 11 · RAG; the long-context doc is doc 19 · Long Context & Memory.
10Post-Training Everywhere
What: RLHF/DPO/LoRA fine-tuning became commodity skills, while the differentiator moved to RLVR for reasoning and "verifier-as-data" — the checker defines the curriculum. Why it matters: when anyone can LoRA a checkpoint, the moat is the reward signal you can compute cheaply and honestly. Math hook: RLVR rewards are checkable predicates (unit tests, exact match), so advantage estimates carry less label noise than preference pairs. Deep dive: doc 17.
05 How to Track a Moving Target
You cannot subscribe to "the frontier." You can subscribe to signal types — each with a different reliability/cost profile:
📄 arXiv & tech reports
Highest information density, highest reading cost. Filter by what changes an equation you already know — a new KV layout, a new routing loss — not by headline model names. Cross-check claims against the appendix tables, not the abstract.
🔄 Provider changelogs
The ground truth for what you can actually ship: pricing shifts, cached-token discounts, context windows, tool-calling APIs. A pricing change is a hardware disclosure in disguise — read doc 07's economics off every new price list.
📊 Benchmarks
Weakest signal, most viral. A benchmark score is a sample from one eval distribution under unknown serving conditions. Trust trends across independent evals more than any single leaderboard jump — and check calibration (ECE-style), not just accuracy.
File each headline under one of the five axes; ask "which bottleneck did this attack?"; hedge every number with its source; re-check this page's claims against primary sources before building on them.
Rewrite production routing on a single-source report (see item 01); compare benchmarks across different serving configurations; assume a demo latency is a served latency; date-stamp knowledge you can't refresh.
06 Mental Models
A news page like this one is a forecast with a freshness half-life. The axes (sparsity, memory, test-time compute, readouts, harnesses) are climate — they change on year scales. Specific models, prices, and benchmark numbers are weather — they change weekly. Invest your reasoning in climate. Lets you reason about: which parts of this doc to memorize (none of the numbers, all of the axes).
Every trend is a bet that one bottleneck (FLOPs, KV bandwidth, decode latency, coordination cost) is now the binding constraint. MoE bets on FLOPs; KV formats bet on bandwidth; test-time compute bets that tokens can buy accuracy; harness standards bet that coordination, not capability, is scarce. Lets you reason about: why two credible announcements can point in opposite directions — they're attacking different walls.
07 Common Misconceptions
"News = product demos." A demo is a sample under curated conditions. Benchmarks ≠ serving reality: latency, cost, calibration, and failure modes only show up under your traffic. The gap between leaderboard and production is exactly docs 07 + 08's economics and memory math.
"A bigger context window means the model remembers everything in it." Window size is an advertisement about capacity, not about fidelity. Attention quality degrades mid-context (doc 06), and every decode step still pays the full KV read (doc 08) — retrieval (doc 11) often beats stuffing.
"Decision models will replace LLMs." The reported JEV class replaces decoding for structured classification and routing — one readout instead of hundreds of generated tokens. Generation, synthesis, and code remain autoregressive territory. Complement, not replacement — and an unverified one at that.
"MoE means the model got smaller." Total parameters grew (671B); only the per-token compute shrank (37B active). You still pay the memory bill for all experts — which is precisely why the KV/memory-format axis (item 04) had to follow it.
"As of this writing" hedging is journalism, not engineering. It's the opposite: production systems that hard-code model behaviors break on the next changelog. Hedged claims + five-axis filing is how you keep an architecture alive across model swaps.